NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Overcoming Catastrophic Forgetting by Generative Regularization

Chen, Patrick H; Wei, Wei; Hsieh, Cho-Jui; Dai, Bo. (January 2021, International Conference on Machine Learning (ICML))
null (Ed.)
Full Text Available
MulCode: A Multiplicative Multi-way Model for Compressing Neural Language Model

Ma, Yukun; Chen, Patrick H; Hsieh, Cho-Jui (October 2019, EMNLP)

Full Text Available
Clustering and Constructing User Coresets to Accelerate Large-scale Top-K Recommender Systems

Jiang, Jyun-Yu; Chen, Patrick H; Hsieh, Cho-Jui; Wang, Wei (January 2020, WWW)

Full Text Available
Learning to Screen for Fast Softmax Inference on Large Vocabulary Neural Networks

Chen, Patrick H; Si, Si; Kumar, Sanjiv; Li, Yang; Hsieh, Cho-Jui (April 2019, International conference on learning representation (ICLR))

Full Text Available
Efficient Contextual Representation Learning With Continuous Outputs

https://doi.org/10.1162/tacl_a_00289

Li, Liunian Harold; Chen, Patrick H.; Hsieh, Cho-Jui; Chang, Kai-Wei (March 2019, Transactions of the Association for Computational Linguistics)

Contextual representation models have achieved great success in improving various downstream natural language processing tasks. However, these language-model-based encoders are difficult to train due to their large parameter size and high computational complexity. By carefully examining the training procedure, we observe that the softmax layer, which predicts a distribution of the target word, often induces significant overhead, especially when the vocabulary size is large. Therefore, we revisit the design of the output layer and consider directly predicting the pre-trained embedding of the target word for a given context. When applied to ELMo, the proposed approach achieves a 4-fold speedup and eliminates 80% trainable parameters while achieving competitive performance on downstream tasks. Further analysis shows that the approach maintains the speed advantage under various settings, even when the sentence encoder is scaled up.
more » « less
Full Text Available
GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking

Chen, Patrick H; Si, Si; Li, Yang; Chelba, Cipirian; Hsieh, Cho-Jui (January 2018, Advances in neural information processing systems)

Full Text Available

Search for: All records